Skip to main content

Overview

This doc will explain tips I learned for rendering the 3D characters, with particular emphasise on how they were implemented into MaO. The same steps tips discussed here should be reusable in other environements as well. If you're looking at implementing a character like this I strongly discourage offloading the work to AI. Developing convincing 3D characters requires proper human input, only you can tell if a character is uncanny, only you can truly understand the impact changing any aspect of the workings of the character has on how it's perceived. An AI may make decent templates and reasonable parameters, but it can't intrinsically feel how we do about near-human characters.

This doc describes my personal learnings and findings when developing these characters, as well as explains the systems currently in place in MaO so they can be improved into the future.

Rendering

Meall an Óige uses ThreeJS to render the character. ThreeJS is supported natively in React and has methods to handle Morph Targets effectively. ThreeJS is usually called in a Canvas. It is important to be mindful of the level you deploy the Canvas, as if you subject it to being destroyed and re-rendered too often there will be a significant impact on performance.

As the characters are used across the app in MaO, I decided to call the canvas at the highest level possible (alongside the main app content in layout.tsx) to assure that events in the app content don't cause re-renders of the canvas unless specifically designed to. In the future (thinking of a standalone Speech-to-Speech page on Abair.ie) it may be desired for a character to only be rendered in one singular page, in which case the canvas/ThreeJS should have extra care put into correctly decomposing and recomposing the renderer.

ThreeJS has a lot of flexibilty in terms of camera angles, light sources, ambient occlusion and more things I dont know about. Meall an Óige required flat lighting with little shadow so a default inbuilt environment from Drei was used, other applications may want fancier visuals.

To improve perform and minimise file size all characters should use KTX2 textures and Draco Compression. The system in Meall an Óige is built to decode KTX2. This site provides a handy way to do Draco Compress and convert a GLB's textures to KTX2 in a browser.

Optimisations for speech​

The characters perform lipsyncing by receiving JSON data containing the timings certain visemes occur at (see the api for more on this). The app then interpolates frame by frame between each viseme shape, keeping track with the timings of the JSON. The lipsyncing is not inherintly tied to the playing of the audio. Instead, the audio and the lipsyncing are simply started at the same time. To account for potential issues like buffering it is recommended to use OnPlaybackStatus or simialar methods to assure lipsyncing actually starts when the audio does, not just when it is called.

To account for the audio-visual delay of real speech, I found offsetting the lipsyncing by -0.075s (lipsyncing happening slightly advanced of the audio, that is) to produce favourable results. In MaO this is handled directly in the lipsyncing funtion, it doesnt have to manually be called 0.075 seconds earlier each time audio is played.

As far as interpolation is concerned, I found using a smoothstep curve (a curve that spends more time at the extremes rather than in the transition state) far superior than a linear curve. The linear curve provided visuals that were, to my eye, "bouncy". The mouth seemed to float around the face from position to position. Smoothstep helped the mouth seem more anchored in place with visemes that felt much more intentional, defined and readable. This was true for all characters tested.

Idle actions​

Idle actions are probably the biggest weakness of the systen in Meall an Óige. The code is built to perform an array of different automated actions when idle or speaking, but the parameters governing things like the frequency and amplitude of actions, and in some cases entire types of actions could do with some polishing. Input from a experienced video game designer could be useful to help tailor the parameters. Nonetheless, I can still share some of my opininions and learnings that have at least helped get the characters part-way there.

Blinking​

Research seems to state that blinking is mostly independent of other facial actions. The research suggests that cognitive tasks increase the blinking rate. With this in mind, the characters in Meall an Óige are programmed to have two blinking intervals. If the character is not speaking, it will blink at intervals of 3-6 seconds. If it is speaking, the interval reduces to 1.5-3.5 seconds.

Gaze​

The characters also look around differently depending on if they are idle or speaking. When the character is speaking, the eyes spend more time not looking at the camera and look further away. When speaking the gaze can still shift, but to a lesser degree and reverts back to center quicker.

Head Bobbing​

The first implemention of head bobbing works much the same way as the gaze does. It is proposed that uniting the Head Bobbing and Gaze into one set of actions would produced better animations as movements would feel more intune with each other. At the moment, the eyes may look left but the head could bob up, down, or right. Testing uniting the actions so the head follows the gaze should be a priority.

Brows​

When not speaking brows occasionally raise or lower. Brow movement is not attatched to any other head movements. When speaking, no automated brow movements occur, but it is encouraged for the component calling the lipsyncing to call actions like performQuestioning/performSad/performSmile depending on the context.

Breathing​

The breathing animation happens irrespective of other actions. It is also affected by wether the character is speaking or not. When idle, the breathes are slower and deeper (higher amplitude, longer interval) and when speaking they have average amplitude but happen more frequently.

Speech Reactions​

Operating on the basis that humans have natural tendencies when begining and ending speaking (things like raising or lowering the eyebrows when starting an utterance that requires reflection, nodding when starting an utterance with enthuasism, relaxing after finishing speaking) some edge cases have been created in MaO. Currently the characters tilt the head forwards slightly on speech start and smile when the speech ends. This doesn't come always come across naturaly so some tailoring would be useful here. Potentially an array of different edge actions should be created so the most appropriate action can be called depening on the context. Otherwise they could just be completely removed.

On-demand Actions​

On-top of the automated actions, there are also on-demand actions that can be triggered anywhere in the app. They can be categorised as either hold actions, or repeating actions. Hold actions (for example emotions) transition to a position, hold for x seconds, then transition back. Repeating actions (like nodding) perform an animation x number of times. Both types of actions can be called with two parameters;

  • intensity (how strong the action should be), defaults to 1.
  • holdDuration or repeatsN. Hold actions take holdDuration, which controls how long in seconds the action hold, repeating actions take repeatsN which governs how many cycles the action performs. These actions can be called dynamically, or by passing through a json train of actions obtained by the API.

Coding for multiple characters​

Obviously we want the system to be able to work with multiple different characters. We also don't want the code to have to be rebuilt for each new character that is added. To account for this, there exists a modelConfig file that converts each characters shape keys into standardised names that the rest of the functions and components can understand. That way if a characters keys are named differently or if its missing certain functionalities another character has, it can still work like the others without major changes. In MaO the config file looks like this (see below), but it may look different depending on what you need the characters to do. As you can see, most Morph targets are optional, so a model doesn't have to include them. If the model doesnt include an shape key used in an automated action, the automated action wont occur.

export const models = {
lion: '/character_models/lion.glb',
spring: '/character_models/spring.glb',
} as const;

export type ModelName = keyof typeof models;
type VisemeMix = {
key: string;
mix: Record<string, number>;
};

type ModelConfig = {
visemeMode: 'mix' | 'direct'; //use direct if there exists a shapeKey matching all 21 visemes GiobGeab outputs by name. use mix if custom mapping has to be used
visemeMixes: VisemeMix[];

//emotions & expressions
smile?: string | Record<string, number>;
sad?: string | Record<string, number>;
worry?: string | Record<string, number>;
shocked?: string | Record<string, number>;
questionFace?: string | Record<string, number>;
squintyFace?: string;

//head movements
headLeft?: string[];
headRight?: string[];
headForward?: string[];
headTiltLeft?: string[];
headTiltRight?: string[];
headBobActions?: Record<string, number>[];

//eyelid movements
eyeSquint?: string;
eyeBlink?: string;
eyeKeys?: string[];
eyeBlinkKeys?: string[];

//eye movements
lookRight?: string | Record<string, number>;
lookLeft?: string | Record<string, number>;
lookUpRight?: string;
lookUpLeft?: string;
lookDownRight?: string;
lookDownLeft?: string;
lookUp?: string | Record<string, number>;
lookDown?: string | Record<string, number>;

//brow movements
browsUp?: string;
browsDown?: string;
browsInnerUp?: string;
browsInnerDown?: string;
passiveBrowActions?: string[];

breatheKeys?: string[];

};
export const MODEL_CONFIGS: Record<ModelName, ModelConfig> = {
lion: {
visemeMode: 'mix',
visemeMixes: [
{ key: "resting", mix: {} },
{ key: "PBM_S", mix: { PBM_S: 1 } },
{ key: "PBM_B", mix: { PBM_B: 1 } },
{ key: "FV_S", mix: { FV_S: 1 } },
{ key: "FV_B", mix: { FV_B: 1 } },
{ key: "S", mix: { TD_S: 1 } },
{ key: "SH", mix: { SH: 1 } },
{ key: "TD_S", mix: { TD_S: 1 } },
{ key: "TD_B", mix: { TD_B: 1 } },
{ key: "KG_S", mix: { TD_S: 1 } },
{ key: "KG_B", mix: { TD_B: 1 } },
{ key: "LSL", mix: { NBL: 1 } },
{ key: "LBL", mix: { NBL: 1 } },
{ key: "NSL", mix: { NBL: 1 } },
{ key: "NBL", mix: { NBL: 1 } },
{ key: "VFO", mix: { VFO: 1 } },
{ key: "VFM", mix: { VFM: 1 } },
{ key: "VFC", mix: { VFC: 1 } },
{ key: "VBM", mix: { VBM: 1 } },
{ key: "VBC", mix: { VBC: 1 } },
{ key: "VBCU", mix: { VFC: 1 } },
],
//emotions
smile: { bigSmile: 0.6 },
sad: 'worried',
questionFace: { eye_squint: 1, headTiltLeft: 0.6 },

//eye movements
lookRight: { eyeRightRight: 0.3, eyeLeftLeft: 0.3, },//ignore mismatching names here, shape keys are mislabled in the file itself
lookLeft: { eyeRightLeft: 0.3, eyeLeftRight: 0.3, }, //ignore mismatching names here, shape keys are mislabled in the file itself
lookUp: { eyeRightUp: 0.3, eyeLeftUp: 0.3, },
lookDown: { eyeRightDown: 0.3, eyeLeftDown: 0.3, },

//head movements
headLeft: ['left_12'], // headLeft: ['left_12', 'left_25'],
headRight: ['right_12'], // headRight: ['right_12', 'right_25'],
headForward: ['forward_10'],
headTiltLeft: ['headTiltLeft'],
headTiltRight: ['headTiltRight'],
headBobActions: [{ forward_10: 0.2 }, { headTiltLeft: 0.3 }, { headTiltRight: 0.3 }],

//eyes
eyeSquint: 'eye_squint',
eyeBlink: 'eye_blink',
eyeBlinkKeys: ['eye_squint', 'eye_blink'], //VERIFY if used
eyeKeys: ['eye_squint', 'eye_blink'],
},
spring: {
visemeMode: 'mix',
visemeMixes: [
{ key: "resting", mix: {} },
{ key: "PBM_B", mix: { PBM_B: 1 } },
{ key: "PBM_S", mix: { PBM_S: 1 } },
{ key: "FV_S", mix: { FV_S: 1 } },
{ key: "FV_B", mix: { FV_B: 1 } },
{ key: "S", mix: { S: 1 } },
{ key: "SH", mix: { SH: 1 } },
{ key: "TD_S", mix: { TD_S: 1 } },
{ key: "TD_B", mix: { TD_B: 1 } },
{ key: "KG_S", mix: { KG_S: 1 } },
{ key: "KG_B", mix: { KG_B: 1 } },
{ key: "LSL", mix: { LSL: 1 } },
{ key: "LBL", mix: { LBL: 1 } },
{ key: "NSL", mix: { NSL: 1 } },
{ key: "NBL", mix: { NBL: 1 } },
{ key: "VFO", mix: { VFO: 1 } },
{ key: "VFM", mix: { VFM: 1 } },
{ key: "VFC", mix: { VFC: 1 } },
{ key: "VBM", mix: { VBM: 1 } },
{ key: "VBC", mix: { VBC: 1 } },
{ key: "VBCU", mix: { VFC: 1 } },
],
//emotions
worry: { browsInnerUp: 1, VFM: 1 }, //not currently used
shocked: { browsUp: 1, VFO: 1 }, //not currently used
smile: { lipsSmile: 1, eyesClosed: 0.1 },
sad: { browsInnerUp: 1, lipsFrown: 1 },
squintyFace: 'squintyFace', //not currently used
questionFace: { browsUp: 0.8, headTiltLeft: 0.5 },

//head movements
headLeft: ['headLeft'],
headRight: ['headRight'],
headTiltLeft: ['headTiltLeft'],
headTiltRight: ['headTiltRight'],
headForward: ['headDown'],
headBobActions: [{ headDown: 0.2 }, { headTiltLeft: 0.2 }, { headTiltRight: 0.2 }],

//eye movements
lookRight: 'eyesRight',
lookLeft: 'eyesLeft',
lookUpRight: 'eyesUpRight',
lookUpLeft: 'eyesUpLeft',
lookDownRight: 'eyesDownRight',
lookDownLeft: 'eyesDownLeft',
lookUp: 'eyesUp',
lookDown: 'eyesDown',

eyeBlink: 'eyesClosed',
eyeBlinkKeys: ['eyesClosed'],
eyeKeys: ['squintyFace', 'eyesClosed'],

// brows
browsUp: 'browsUp',
browsDown: 'browsDown',
browsInnerUp: 'browsInnerUp',
browsInnerDown: 'browsInnerDown', //not currently used

passiveBrowActions: ['browsDown', 'browsUp'],

breatheKeys: ['breatheIn'],
}
};

export function buildVisemeLookup(config: typeof MODEL_CONFIGS[ModelName]): Record<string, Record<string, number>> | null {
if (config.visemeMode === 'direct' || !config.visemeMixes) return null;
return Object.fromEntries(
config.visemeMixes.map(({ key, mix }) => [key, mix])
);
}